Skip to content

Add DSV4 GB200 Dynamo+SGLang AgentX recipes without W4A4 MegaMoE - #3629

Merged
cquil11 merged 7 commits into
mainfrom
dsv4-fp4-gb200-dyn-sgl-no-megamoe
Oct 2, 2026
Merged

cquil11 merged 7 commits into
mainfrom
dsv4-fp4-gb200-dyn-sgl-no-megamoe

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Add DSV4-Pro GB200 Dynamo+SGLang AgentX recipes using DSpark block size 6 and HiCache on lmsysorg/sglang:nightly-dev-20260916-c9a8fba9.

Coverage: TP8 aggregate at C1/C4; 1P1D DEP8/DEP16 at C64/C128; 1P1D DEP16/DEP32 at C256; and 2P1D DEP16/DEP32 at C768/C1024/C1280.

The disaggregated recipes omit enable-w4a4-mxfp4-megamoe, retain the MegaMoE all-to-all backend and FP4 indexer, and use decode request caps of 512/1024 for DEP16/DEP32. C128 prefill uses mem-fraction-static: 0.80.

nvpohanh and others added 4 commits October 1, 2026 00:56
Remove enable-w4a4-mxfp4-megamoe from every prefill and decode worker in
the six GB200 Dynamo+SGLang AgentX disaggregated recipes. The MegaMoE
all-to-all backend and FP4 indexer settings are unchanged.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@nvpohanh nvpohanh added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Oct 1, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked the PR's core claim — that enable-w4a4-mxfp4-megamoe was removed from all prefill/decode workers in the six disaggregated recipes while moe-a2a-backend: megamoe stays — and confirmed no enable-w4a4-mxfp4-megamoe key remains anywhere in the new recipe files, and container.image matches the master-config image field consistently across all eight new scenarios.

Extended reasoning...

The diff adds 7 new srt-slurm recipe YAMLs plus matching master-config entries, a runners.yaml model-path mapping, a lanes.py policy-table entry, and an append-only changelog addition for a new DeepSeek-V4-Pro-0813 GB200 FP4 AgentX workload; no auth/crypto/permission code is touched, only benchmark config and launch policy tables. A confirmed finding already flags a real routing problem in runners.yaml (new model lacks a numa1-local alias, so it is not visible for node-local staging as the changelog implies), which is sufficient reason a human should look before merge.

Comment thread inferencex-e2e/configs/runners.yaml
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

@Ankur-singh

Copy link
Copy Markdown
Collaborator

/use 36833446190

Reuse the passing full sweep for this PR: attempt 1, tested commit 2a58623b5a441f1b5c190ba9666a70881301d1d3.

中文

复用本 PR 已通过的完整 sweep:attempt 1,测试 commit 2a58623b5a441f1b5c190ba9666a70881301d1d3。

@Ankur-singh

Ankur-singh commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Assessed 2a58623b5a441f1b5c190ba9666a70881301d1d3 against the current checklist at 15b01bb5. CI passed; Run Sweep 36833446190, attempt 1 executed 8 performance points and 6 eval leaves on this head. Performance; accuracy: GSM8K em_strict 97.57–97.95%, n_eff=1319 per eval, same SGLang image.
  • Chat-completions replay; automatic SGLANG_SIMULATE_ACC_LEN=3.77, match-expected, real-draft-token, from DSpark Golden AL, thinking_on, draft length 6. Eval-only worker environments contain no synthetic acceptance settings.
  • Bundled DSpark mtp.0/1/2 heads from deepseek-ai/DeepSeek-V4-Pro-0813. Checkpoint metadata: mixed native FP4 experts, FP8 projections and BF16/F32 tensors. Pinned SGLang loader preserves the shipped mixed precision with its default FP8 wo_a conversion to BF16; effective recipes add no draft quantization, substitution or patch. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is unset in logged worker environments and defaults to false.
  • Engine-first item verified under the SA admin decision merged in #3667. Current verifier Check 6(b) treats Dynamo as a deployment layer: dynamo-vllm and dynamo-sglang are open-source engine entries, and a dynamo-sglang submission does not require a preceding plain sglang entry for the same model/SKU. This PR adds only dynamo-sglang entries for dsv4 / cluster:gb200-nv, so the ordering check passes.
  • Upstream lmsysorg/sglang:nightly-dev-20260916-c9a8fba9; official Dynamo 1.5.0.dev20260910 wheel. No engine patches. Upstream single-node recipe link and append-only: true checks are N/A: every added recipe spans physical nodes, and these are regular changelog entries.
  • P90 throughput/E2EL: 6 measured frontier points from 8 results on one GB200/Dynamo-SGLang/FP4 curve; frontier concurrencies 4, 64, 256, 768, 1024, 1280. No Pareto bypass needed. All power audits are invalid (REQUIRE_POWER=0); no power/energy claim. Changelog's node-local staging wording conflicts with the shared Lustre mapping.
  • Reuse selected by authorized /use 36833446190 comment. PR currently has merge conflicts; synchronization needs a fresh assessment. This checklist is separate from formal approval and the official bot verdict.

Signed: Ankur-singh

@Ankur-singh Ankur-singh left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the benchmark and eval evidence for 2a58623b5a441f1b5c190ba9666a70881301d1d3. CODEOWNER checklist. The engine-first item remains unchecked pending core-maintainer resolution. Merge conflicts also remain.

中文

已审阅 2a58623b5a441f1b5c190ba9666a70881301d1d3 的 benchmark 和 eval 证据。CODEOWNER checklist。engine-first 项仍未勾选,等待核心维护者确认;merge conflicts 也仍待解决。

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

✅✅✅ Verdict: PASS ✅✅✅

Assessed pinned head 2a58623b5a441f1b5c190ba9666a70881301d1d3; the PR tip has not moved. Note: the PR currently has merge conflicts, so a re-sync commit would need a fresh assessment.

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — @ankur-singh is listed as an owner of inferencex-e2e/configs/nvidia-master.yaml, the only specifically owned changed path. All other changed files fall under the * catch-all.

✅ Check 1 (Passing sweep on in-PR commit): PASS — On pinned head 2a58623b, Run Sweep 36833446190 (attempt 1) shows all 8 multi-node agentic / and all 6 multi-node agentic eval / check-runs as success. None were skipped.

✅ Check 2 (Evals pass): PASS — eval_results_all holds 6 GSM8K results with em_strict 0.9757–0.9795 and n_eff=1319 each (conc 64/128/256/768/1024/1280). The eval jobs ran in the same run on lmsysorg/sglang:nightly-dev-20260916-c9a8fba9.

➖ Check 3 (Upstream recipe): N/A — disaggregated/multi-node submission; the recipe-link requirement applies to single-node recipes only. All 8 recipes are under multi_node/srt-slurm-recipes/ and both entries are multinode: true.

✅ Check 4 (Reuse command): PASS — /use 36833446190 was posted by Ankur-singh (COLLABORATOR).

✅ Check 5 (Latest checklist template): PASS — All 17 items in the current PR_REVIEW_CHECKLIST.md template, including the nested upstream-PR item, are present and checked.

✅ Check 6 (Upstream images / engine-first): PASS — Both new entries use upstream lmsysorg/sglang:nightly-dev-20260916-c9a8fba9 with framework: dynamo-sglang. These are open-source engine entries only, so the ordering rule does not apply.

✅ Check 7 (No deprecated models/scenarios): PASS — dsv4 agentic coding is active in MODELS.md. Only 8k1k/1k1k are deprecated for dsv4.

✅ Check 8 (No architecture hacks / event publication): PASS — There are no --hf-overrides, model-override args or model-file edits. publish_events_and_metrics is a TRT-LLM setting that SGLang does not support, and no recipe or inherited base sets it (the recipes are standalone, with no base).

✅ Check 9 (Spec-decode via chat template): PASS — The logged aiperf command uses --endpoint /v1/chat/completions --endpoint-type chat, from the conc64 server-log artifact.

✅ Check 10 (No engine patches): PASS — The diff has no .patch/git apply/sed -i/heredoc rewrites, no engine pip install and no file overwrites. The lanes.py/runners.yaml changes are harness routing and model-path mapping only.

✅ Check 11 (Agentic golden AL): PASS — synthetic_acceptance.py injects SGLANG_SIMULATE_ACC_LEN=3.77, match-expected and real-draft-token into the prefill and decode envs. The logged config.yaml and fingerprints confirm this. The value matches dsv4-pro-0813-dspark.yaml thinking_on at DSpark block size 6 (3.77).

➖ Check 12 (Append-only): N/A — Neither new perf-changelog.yaml entry sets append-only: true.

✅ Check 13 (Draft runs as shipped): PASS — The draft is the bundled DSpark head of deepseek-ai/DeepSeek-V4-Pro-0813, loaded through the pinned SGLang c9a8fba9 default path (default FP8 wo_a→BF16 dequant). The recipes add no draft quantization, dtype override, draft path or DSpark env knobs. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is absent from the recipes and logged worker envs, and defaults to EnvBool(False) (environ.py:985).

✅ Check 14 (Pareto coverage): PASS — One curve: dsv4 / agentic-coding / GB200 dynamo-sglang / fp4 / P90 E2EL / run 36833446190 / lmsysorg/sglang:nightly-dev-20260916-c9a8fba9. It has 8 valid measured points from the bmk_agentic_* artifacts and 6/5 frontier points (conc 4, 64, 256, 768, 1024, 1280), reproduced with infx.workflows.pareto_coverage.

Assessed commit: 2a58623b5a441f1b5c190ba9666a70881301d1d3.

shared_run_root=(
Match(any_of("minimaxm3", "kimik3", "qwen3.5", "glm5.2")),
Match(any_of("dsv4"), frameworks=any_of("dynamo-vllm")),
Match(any_of("dsv4"), frameworks=any_of("dynamo-sglang"), agentic=True),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@adibarra can you take a look since change to infx

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@cquil11 cquil11 changed the title Add DSV4 GB200 Dynamo+SGLang AgentX configs without W4A4 MegaMoE / 添加不启用 W4A4 MegaMoE 的 DSV4 GB200 Dynamo+SGLang AgentX 配置 Add DSV4 GB200 Dynamo+SGLang AgentX recipes without W4A4 MegaMoE Oct 2, 2026
Preserve main changelog entries and append the GB200 DSV4 contribution.
@cquil11
cquil11 merged commit 509bd45 into main Oct 2, 2026
25 checks passed
@cquil11
cquil11 deleted the dsv4-fp4-gb200-dyn-sgl-no-megamoe branch October 2, 2026 15:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

4 participants